Journal of Vision
● Association for Research in Vision and Ophthalmology (ARVO)
Preprints posted in the last 90 days, ranked by how well they match Journal of Vision's content profile, based on 110 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.
de Jong, J.; Sergent, C.; Wexler, M.
Show abstract
The temporal resolution of vision is seriously limited. However, the response to a very brief flash, called the impulse response, is already quite sluggish at the earliest stages of vision, potentially obscuring the true temporal resolution of the rest of the visual system. Faster monitors that produce briefer flashes are subject to diminishing returns because, by definition, they cannot elicit responses that are any briefer than the impulse response. Here, taking inspiration from previous attempts, we develop a novel technique for presenting flashes that elicit 'briefer-than-brief' visual responses. Using a simple deconvolution technique, we reverse-engineer the visual response and estimate the form that the stimulus should take to elicit the response that a faster visual system would produce to a normal flash. Using psychophysics on human observers, we demonstrate that these 'briefer-than-brief' (BTB) flashes partially bypass the temporal limits presumably imposed by the early visual system using two paradigms: one that requires temporal segregation and one that requires temporal integration of sequential flashes. With BTB flashes, human observers successfully isolated two successive flashes at shorter intervals than with conventional flashes, improving temporal resolution by around 16%. We found that BTB stimuli not only improved temporal resolution, but also induced poorer performance on tasks requiring temporal integration, suggesting that the visual responses elicited by BTB flashes overlap less in time due to their briefer duration. In sum, our findings suggest that, using reverse-engineered stimuli, we can alleviate a temporal bottleneck that probably originates from the earliest stages of vision. In doing so, we allow higher visual areas to operate at a higher temporal resolution than previously thought possible.
Ollikka, N.; Bergstrom, A.; Kilpelainen, M.; Deny, S.
Show abstract
Mounting evidence suggests that recurrent processes in the visual system play a critical role during challenging recognition tasks. Backward masking techniques have traditionally been used as a non-invasive method for studying recurrent processes: A mask follows the target image, presumably disrupting ongoing processes. However, these techniques have the limitation that they do not allow the identification of the stage of the visual system at which critical recurrent processes are taking place. Here, leveraging advances in texture synthesis via deep networks, and the approximate correspondence between stages of the visual system and layers of deep networks, we develop a novel psychophysics paradigm where masks with textures targeting different stages of the visual system follow the presentation of challenging images. In a series of experiments, we present objects to human subjects either for a short duration or in unusual poses, followed by a textured mask either designed to only target the early visual system, or the entire visual system. We find that both texture types equally affect recognition abilities, suggesting that recurrent processes in or towards early stages of the visual system are already recruited for these recognition tasks.
Prahalad, K. S.; Poletti, M.
Show abstract
Fixation is often treated as a period of stable visual processing. Yet, fixation is often punctuated by frequent microsaccades that occur during tasks involving complex foveal stimuli. These small eye movements are preceded by changes in visual sensitivity, both at the upcoming movement goal and at the currently fixated location. However, previous work has largely focused on isolated stimuli, leaving unclear whether pre-microsaccadic modulations reflect changes in sensitivity alone or also alter the spatial interactions that govern object recognition. Visual crowding provides a direct test of this question because it depends on the integration and segregation of nearby features and constrains recognition even within the foveola. Using high-precision Dual Purkinje Image eye tracking with retinally contingent stimulus delivery, we measured acuity and crowding thresholds at the preferred locus of fixation (PLF), the starting point of the impending gaze shift, while observers either maintained fixation or prepared to execute a microsaccade to a cued location. Unflanked acuity at the PLF remained stable across conditions. In contrast, crowding strength increased during the pre-microsaccadic interval, indicating an expansion of the foveal crowding zone. These results show that microsaccade preparation alters spatial integration at the starting point of the movement, increasing crowding even when sensitivity to isolated stimuli remains unchanged. Thus, microsaccades reshape foveal vision not only by modulating visual discrimination at the movement goal, but also by changing how nearby features are integrated and segregated before the eyes move. Significance StatementVision is often assumed to be most stable when gaze is fixed. Yet the eyes are never truly still, and the brain continually prepares small movements that shape perception before they occur. This study shows that such preparation changes how visual information is organized at the very center of gaze. Upcoming eye movements do not simply alter sensitivity to isolated objects; instead, they change how nearby features are integrated. These findings reveal that fine spatial vision and object recognition depends not only on what falls on the retina or on subsequent cortical processing but also on what the eyes are preparing to do next.
Meidan, R. Y.; Bonneh, Y. S.
Show abstract
Visual discomfort (VD) is influenced by both spatial structure and chromatic context. Striped patterns are well-established triggers of discomfort and autonomic responses. In previous work, we showed that higher spatial frequencies and larger patterned areas elicit stronger pupillary constriction and greater discomfort, and that individuals with higher overall discomfort show shallower maximum constriction. The present study examined whether similar relationships appear when spatial structure is held constant, measured background luminance is kept within a narrow range, and the chromatic background varies. Participants viewed black horizontal stripes on 12 near-isoluminant colored backgrounds. The CIE76 color difference ({Delta}E) ranged from 36 to 112 and was indexed as the CIELAB distance from the black stripe pattern. Pupil size was continuously recorded and discomfort ratings were collected after each trial. Across the colored backgrounds, more uncomfortable stimuli evoked stronger pupil constriction, even though luminance was held nearly constant. As in our spatial-frequency study, this stimulus-level increase did not translate into stronger constriction among observers reporting higher overall discomfort: participants who rated the stimuli as more uncomfortable overall showed shallower constriction. This pattern was captured by the maximum-constriction response, which differentiated high-from low-discomfort observers and was significantly associated with individual discomfort ratings. Together with our previous findings, the results suggest that pupil responses scale along the tested stimulus axis, whereas individuals reporting greater visual discomfort exhibit less pronounced maximum constriction. HighlightsO_LIChromatic background modulated discomfort and pupil responses at similar luminance. C_LIO_LIAcross backgrounds, higher discomfort ratings tracked stronger pupil constriction. C_LIO_LIAcross observers, higher mean discomfort tracked weaker pupil constriction. C_LIO_LIThis two-level dissociation recurs across spatial and chromatic manipulations. C_LIO_LIPupillometry may complement subjective reports of visual discomfort. C_LI
Peterzell, D. H.; Arrighi, R.; Di Cesare, C.; Gurioli, M.; Farini, a.; Grasso, P. A.
Show abstract
Numerosity adaptation (the underestimation of number after exposure to a numerous adaptor) is reduced when adaptor and test differ in color, suggesting that the numerosity system parses items into color-defined categories. Here we ask whether this chromatic selectivity is organized into multiple narrowly tuned chromatic channels, and whether its expression depends on individual chromatic sensitivity. Twenty observers (aged 22-61) completed two psychophysical tasks. First, chromatic discrimination was measured for five hues spaced in 5{degrees} CIE L*a*b* steps ({Delta}H = 0{degrees}, 5{degrees}, 10{degrees}, 15{degrees}, 20{degrees}) from a red reference (LCh: 54, 118, 38), yielding an individual just-noticeable difference (JND). Second, numerosity adaptation was measured across the same five chromatic distances between a 48-dot adaptor and the test. Observers with superior discrimination (JND < 2.5{degrees}) showed robust chromatic tuning, adaptation declining as the test moved away from the adaptor hue, whereas poorer discriminators showed none. Using an interindividual-covariance / factor-analytic approach, we found that adaptation strengths at neighboring chromatic distances were highly correlated and fell off with chromatic separation. Principal component analysis extracted two factors, one loading on the larger chromatic distances and one on the smaller; under oblique (promax) rotation the two factors were substantially correlated (r = .66), implying at least two dissociable but overlapping chromatically tuned mechanisms. These results suggest that numerosity adaptation is mediated by multiple, comparatively narrow chromatic channels, resembling the higher-order color mechanisms inferred from color scaling, SSVEP, and fMRI, rather than the two early cardinal axes (L-M, S-(L+M)).
Collins, T.
Show abstract
Mental representations are the explanatory construct of the cognitive sciences, but there is no widely accepted characterization of how they cause behavior. Visual representational geometry can be quantified by similarity scores, but almost all methods require an explicit judgment. To examine how representations cause behavior by varying task demands, observers must perform different tasks while continuously reporting similarity, leading to dual-task interference. This study develops scanpaths as an implicit similarity measure, and uses representational similarity analysis to validate it. Observers searched for a target; fixations on distractors may reveal similarity. Similarity was also quantified by an odd-one-out task in the same participants, and ratings from different participants (Jiang et al. 2022). Representational geometries between tasks correlated. A generative model predicted first fixations in novel data. This double validation of the scanpath method opens the door to examining the causality of representations by determining if and how they vary with task demands.
Mendez, A. H.; Otero-Millan, J.; de la Malla, C.; Lopez-Moliner, J.
Show abstract
Rigorously tracking eye and head behavior in space is key to building realistic models of the stimulus that reaches our retina. The motion structure of this stimulus or retinal flow - the substrate for self and object motion processing - is created by the relative movement of the eyes with respect to the world. Characterizing this stimulus requires tracking the eyes three degrees of freedom in the head and the heads six degrees of freedom in the world. While vertical and horizontal eye rotations have been described during locomotion in the context of gaze stabilization (Moore et al, 2001), the component around the line of sight - torsion - has remained difficult to quantify, and how all three rotational components jointly contribute to retinal flow during self-motion remains largely unexplored. Here, we leveraged head-mounted technology to estimate eye torsion in ten subjects as they walked towards a distant target in a fast and slow condition (from 14 to 4 meters away from the target, see Fig. 1A). More specifically, we combined automatic feature tracking with gaze-constrained simulations of eye rotations and camera projection to recover torsion from image data. We then estimated flow curl in head and retina centered frames in two scenarios: torsion as estimated from our data and with no torsion. We show that the eyes torsional component compensates for the roll component of heads angular displacement, altering the incoming visual flow in ways that are relevant for the extraction of self-motion parameters from retinal flow. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/743586v1_fig1.gif" ALT="Figure 1"> View larger version (38K): org.highwire.dtl.DTLVardef@391b9eorg.highwire.dtl.DTLVardef@1444510org.highwire.dtl.DTLVardef@1121e16org.highwire.dtl.DTLVardef@754b6a_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig 1.C_FLOATNO A. Top. Custom-made head-mounted device combining the Neon eye tracker (Pupil Labs), an RGB camera and a dimmable light. Bottom. Four 3D frames of reference (FoR) are relevant for this study, two static (world and locomotion) and two subject centered (head and eye). The Z axis of the locomotion, head and eye FoRs are approximately aligned throughout the trial. For the locomotion FoR the Z axis is fixed in the world and points forward (towards the target). The heads Z axis moves with the head but - as subjects are fixating a target along their path -, it also points approximately forward. The eyes Z axis also moves with the head and its exact forward orientation will depend on compensatory eye movements. B. Left. Blue dots represent the Z component of the heads orientation vector (on the locomotion frame) on the X axis, and the sum of all three components on the Y axis; for each frame for all corpus data. Blue contour is the 75th percentile 2D density distribution of the blue dots. Red and violet contours represent the 75th percentile for the X and Y components of head orientation, respectively. Right. Same logic but applied to the heads velocity vector. C. Left. Two examples showing the mean rotation of iris features over the course of a slow (top) and fast (bottom) trial. Colored lines show each of the 561 simulated cameras for a given scenario (one color per scenario); the black line shows the camera from the empirical data. Right. Trial-level mean fit score of each scenario with the empirical data is represented as a function of each subjects fitted gain. 20 dots represent 10 subjects x 2 trials. On the rightmost column, all values are aligned vertically to show the mean fit score across trials for the three scenarios. Size indicates the 75th percentile of head z component for each trial. C_FIG
Penaloza, B.; Maniglia, M.; Munneke, J.; Green, C. S.; Seitz, A.
Show abstract
Purpose: To evaluate the feasibility, validity, and scalability of PLFest, an open-source, Unity-based, cross-platform application designed for standardized, multi-site visual and cognitive assessment and training. Methods: Two hundred sixty participants (mean age = 23 years) were recruited across four university sites in the United States. Participants completed a battery of five visual assessments administered through PLFest, including visual acuity, contrast sensitivity, spatial frequency cutoff, contrast sensitivity at spatial-frequency cutoff, and visual search. Five cognitive assessments measuring visuospatial working memory, verbal working memory, fluid reasoning, inhibitory control, and selective attention were also administered. Descriptive statistics and performance distributions were examined and compared with normative data. Results: Visual acuity and contrast sensitivity measures closely matched previously reported normative values obtained using established clinical and psychophysical methods. Spatial frequency cutoff and visual search tasks produced stable threshold estimates while showing substantial inter-individual variability. Performance across all cognitive assessments was consistent with published validation studies of the corresponding tasks. Across the full battery, adaptive procedures demonstrated reliable convergence and generated well-distributed performance measures without evidence of substantial floor or ceiling effects. Importantly, these findings were observed across four geographically distributed testing sites using standardized consumer-grade tablet hardware. Conclusions: PLFest provides reliable and scalable assessment of visual and cognitive function using portable consumer devices. The platform supports standardized data collection across distributed research settings while maintaining performance characteristics consistent with established laboratory and clinical benchmarks. These findings support the use of PLFest as a reliable framework for large-scale studies of vision and cognition. Translational Relevance: By reducing dependence on specialized laboratory infrastructure and trained personnel, PLFest may facilitate broader access to visual and cognitive assessment, enabling large-scale research, screening, and future rehabilitation applications.
Casco-Rodriguez, J.; Hong, F.; Brainard, D. H.; Feather, J.; Lipshutz, D.
Show abstract
Representations of the same physical stimulus vary between individuals. Characterizing individual differences has practical implications, but is challenging because these representations are not directly observable. Given a model of how representations vary within a population, we propose a Bayesian adaptive procedure for estimating an individual observer's representation from a series of targeted perceptual discrimination judgments. A key component of our approach is using Fisher information to identify stimulus distortions that efficiently differentiate observers in the population. As a proof of concept, we focus on individual differences in color perception and simulate observers with cone fundamentals drawn from an individual colorimetric observer model. We demonstrate that our approach can recover key aspects of a sampled observer's cone fundamentals using simulated three-alternative forced-choice oddity judgments with approximately 500 trials, corresponding to an experimental duration of approximately one hour. Our Bayesian adaptive framework provides a promising and generalizable approach to efficiently link behavioral measurements to individual differences in sensory representations.
Faul, F.; Nuthmann, A.
Show abstract
Current debates regarding the relative contribution of saliency versus semantics to gaze control often rely on comparing the predictive power of saliency and meaning maps. We argue that such indirect, global approaches are fundamentally limited because fixations arise from heterogeneous, local causes that are conflated in whole-scene comparisons. To substantiate this claim, we used a direct method where participants explicitly identified the reasons for fixation at specific clusters of high fixation density, distinguishing between low-level saliency and various semantic categories, as well as the most important one. The obtained judgments revealed that multiple factors contribute simultaneously to gaze control. Although their influence varied across fixation clusters, semantics generally dominated saliency. Notably, abstract semantic categories, particularly "unknown/unusual," proved important, highlighting the role of prior knowledge and novelty besides personal relevance in guiding attention. To interpret these findings in the context of existing models, we propose a framework distinguishing between processes highlighting interesting locations in the image from a sampling strategy translating this information into scanpaths. Within this framework, classic saliency and meaning maps are viewed as restricted inputs to the strategy, whereas deep learning-based models (e.g., DeepGaze IIE) are more general and may also implicitly encode aspects of the strategy itself. Consistent with this, we found that the predictive performance of DeepGaze IIE varied less significantly with the specific reasons for fixation than that of classic saliency and meaning map approaches.
Castellotti, S.; Blini, E.; Del Viva, M. M.; Zaidi, Q.
Show abstract
Artists and photographers use diagonals to convey dynamism, tension and instability in naturalistic and abstract images. However, the oblique effect in visual perception refers to a reduced sensitivity to oblique compared to cardinal (vertical/horizontal) orientations, linked to anisotropy in the distribution of preferred orientations of primary visual cortex neurons, which reflect the statistical distribution of edge orientations in natural images. Distal diagonals generically project onto oblique orientations on the retina, creating a conundrum between their perceptual salience in art and their reduced sensitivity in vision. We conjectured that global spatial configurations could shape perceived dynamism overriding local orientations. Inspired by van Doesburgs 1929 Arithmetical Composition, we created 60 distinct linear configurations of four identically-shaped rhombi each, in which global and local-edge orientations were independently manipulated to be cardinal or oblique, and measured their perceived dynamism with behavioral, psychophysical, and psychophysiological methods. First, 250 observers rated oblique configurations as more dynamic, tridimensional, and unstable than cardinal configurations, with wedge shapes and element spacing enhancing dynamism. Second, pairwise comparisons (50 observers) yielded a perceptual dynamism scale that confirmed the dominant role of global orientations. Third, pupillometry (35 observers) showed that pupil dilation correlates systematically with subjective dynamism ratings. Perceived dynamism in static images is thus driven primarily by the orientation and shape of spatial configurations, indicating the importance of estimating orientations and shapes of multielement objects beyond decoding local orientations. By linking visual structure to emotional perceptual experience, this investigation provides empirically grounded principles for creating dynamic percepts in visual art. SIGNIFICANCE STATEMENTDiagonals have been used since Michelangelo to create dynamic compositions in representative and abstract art and photography, but paradoxically humans have been shown to be less sensitive to oblique orientations like those projected on the retina by diagonal edges. We resolve this conflict by distinguishing between global configurations versus local edges and showing through behavioral, psychophysical, and psychophysiological experiments that the global orientation and shape of a multi-element configuration has a much larger effect on perceived dynamism than do the edge orientations of its elements. Neural estimation of orientations and shapes of multi-element configurations is thus essential for image understanding. Our results provide perceptually grounded bases for using diagonal configurations with wedge shapes to create dynamic percepts in visual art.
Spitschan, M.
Show abstract
PurposePupil diameter in daily life depends on both the light reaching the eye and the observers age, but established prediction formulas require laboratory quantities that are rarely measured in natural environments. We developed a compact age-corrected model that predicts pupil diameter from melanopic equivalent daylight illuminance (mEDI). MethodsWe used an existing field dataset in which binocular pupil diameter and near-corneal spectral irradiance were recorded while 83 adults aged 18-87 years moved through indoor and outdoor environments. The analysis included 10,082 valid paired observations. We fitted a bounded sigmoid relating pupil diameter to mEDI and age, with each participant given equal influence, and assessed prediction in participants excluded from model fitting. Performance was compared with simpler models, a flexible generalised additive model (GAM), and Watson-Yellott predictions based on assumed field geometry. ResultsPupil diameter decreased smoothly as mEDI increased. Age primarily reduced the difference between pupils in dim and bright conditions, by 0.768 mm per decade, while the predicted bright-light diameter changed little with age. In held-out participants, the bounded model had a participant-balanced root mean squared error (RMSE) of 0.630 mm and mean absolute error of 0.537 mm. The GAM had a slightly lower point-estimate RMSE of 0.610 mm, but the difference was small and uncertain. The bounded model outperformed the tested log-linear, reduced, age-only, and Watson-Yellott alternatives. ConclusionAge and mEDI are sufficient to provide useful population-average pupil predictions across the observed adult age and real-world light range. The model is transparent, physiologically bounded, and nearly as accurate as a flexible GAM, but predictions approaching darkness remain uncertain because valid mEDI measurements were not available in that range. Key pointsO_LIA compact equation predicts population-average pupil diameter from age and mEDI alone. C_LIO_LIAge mainly compresses the pupils response range by reducing pupil diameter under dimmer conditions. C_LIO_LIPrediction error in unseen participants was close to that of a flexible GAM, without requiring a fitted smooth object. C_LIO_LIThe model is intended for the observed adult age and field-light range, not for extrapolation into darkness. C_LI
Michaud, C.; Baures, R.; Soler, V.; Trotter, Y.; Vattier, V.; Rosito, M.; Peyrin, C.; Cottereau, B. R.
Show abstract
Multiple object tracking (MOT) is a core function of dynamic visual attention that relies on the ability to simultaneously monitor several moving objects. Although MOT performance is known to decline with age, and to depend on efficient oculomotor strategies, how these processes interact across the adult lifespan and under degraded visual input remains poorly understood. Here, we examined the effects of aging on MOT under normal and gaze-contingent viewing conditions simulating central and peripheral visual field loss. Sixty participants aged 20-80 years completed a MOT task while eye movements were recorded, enabling characterization of performance and oculomotor behavior across five viewing conditions. Behavioral results revealed a continuous decline in tracking performance across adulthood, indicating a graded rather than categorical effect of age. Performance was strongly reduced by visual-field restrictions, with the largest impairments under central vision occlusion. Eye-tracking analyses showed that better performance was associated with greater reliance on centroid-based gaze strategies, consistent with distributed monitoring of target configurations. Critically, older adults relied more on focal, target-based tracking under conditions simulating peripheral vision loss, and less on centroid-based strategies; this shift was associated with poorer performance. In contrast, oculomotor behavior during full-field viewing was largely preserved across age. Together, these findings suggest that aging affects multiple object tracking through combined sensory, attentional, and oculomotor mechanisms. Beyond a reduction in capacity, age-related decline also reflects systematic changes in visual sampling strategies during dynamic tracking.
Au, D. D.; Melander, J. B.; Weddington, J. C.; Faragalla, Y.; Alaoui, Z.; Liu, S.; Xu, Q.; Baccus, S. A.
Show abstract
BackgroundMice make substantial eye movements during head-fixed visual stimulation, and uncorrected gaze shifts corrupt receptive field measurements and confound stimulus-response relationships. Corneal-reflection video oculography in rodents has provided the methodological foundation for calibrated angular gaze tracking since Stahl (2004) but the calibration procedures used by existing methods -- physical camera rotation, motorized stages, behavioral tasks, or precisely co-aligned dual cameras -- have limited their adoption in many mouse neuroscience laboratories. Most studies instead use uncalibrated pupil tracking, deep learning pose estimation that returns pixel coordinates without angular calibration, or learned shifter networks that lack independent validation. New methodWe present an open-source corneal-reflection eye tracking system for head-fixed mice with two methodological contributions. First, a geometric model recovers gaze in calibrated angular units from the pixel displacements of the pupil and corneal reflections, using the known 3D positions of multiple fiducial LEDs as the source of angular scale. The model requires no estimate of Rp, the per-animal eye-geometry parameter that earlier corneal-reflection methods determine through physical calibration. Second, a self-calibration procedure exploits the redundancy of multiple stationary fiducial LEDs: each LED produces an independent gaze estimate from the same geometric model, and a single residual calibration parameter is determined by minimizing the disagreement between per-LED estimates. This replaces the physical camera-rotation calibrations of earlier video oculography (Sakatani and Isa, 2004, 2007; van Alphen et al., 2013; Kretschmer et al., 2017), the motorized stages of Zoccolan et al. (2010), and the precision dual-camera alignment of Payne and Raymond (2017) with a software operation that requires no moving parts, no behavioral task, and no per-animal procedure. The system provides three interactive GUI stages: (1) pupil and LED detection via Difference-of-Gaussians filtering, (2) 3D geometry definition, and (3) gaze angle computation with blink detection, fiducial correction, and manual curation. ResultsValidation against a rotary-encoder-controlled artificial eye demonstrated mean absolute errors below 1{degrees} across all four fiducial LEDs over the {+/-}20{degrees} working range of mouse eye movements, with Pearson correlations exceeding 0.998 between our methods estimation and encoder ground truth. The self-calibration reduced inter-LED disagreement by a factor of 4-6 in mouse recordings. Gaze-corrected stimulus reconstruction applied to Neuropixels recordings from mouse V1 produced qualitatively sharper receptive field estimates with improved signal-to-noise ratios. Comparison with existing methodsOur method is the first multi-LED, single-camera, fully software-calibrated corneal-reflection eye tracker for mice and includes an integrated open-source pipeline for detection, calibration, blink handling, and artifact correction. The multi-LED redundancy doubles as an internal consistency check -- if two LEDs disagree on gaze direction, the calibration is wrong -- providing a guarantee that learned approaches relying on neural-data-derived correction cannot offer. ConclusionsOur method makes calibrated corneal-reflection eye tracking accessible to non-specialist mouse laboratories using consumer-grade hardware ([~] $2,000-2,700 USD), eliminates the per-animal calibration procedures of earlier methods, and is validated by two independent ground truths at both the absolute angular (artificial eye) and functional (V1 receptive fields) levels. HighlightsO_LIOpen-source corneal-reflection eye tracking for head-fixed mice using a single camera and multiple stationary fiducial LEDs. C_LIO_LIGeometric gaze model derives angular scale from LED positions, eliminating per-animal eye-geometry calibration. C_LIO_LISelf-calibration via multi-LED redundancy replaces physical camera rotation, motorized stages, and dual-camera precision alignment. C_LIO_LIValidated to sub-degree accuracy against a rotary-encoder ground truth across the {+/-} 20{degrees} range of mouse eye movements. C_LIO_LIGaze correction produces sharper V1 receptive field estimates in Neuropixels recordings. C_LI
Jörges, B.; Kim, J.-J.; Harris, L. R.
Show abstract
Continuous Psychophysics, which couples a continuous stimulus with a continuous response, is a promising tool to break out of the confines of traditional designs based on discrete trials. In this pre-registered study, we explore to what extent this paradigm is useful in the study of multisensory integration. We expand on Tonelli et al.s (2025) seminal study by additionally examining the role of eye-movements, using a Kalman filter to estimate the sensory noise underlying behavioral tracking parameters and employing a virtual reality set-up. We immersed two cohorts of participants (n = 30 each) in a virtual meadow environment and asked them to continuously track a drone (Experiment 1) or a swarm of flies (Experiment 2) with a controller, while simultaneously recording their eye movements. We manipulated the reliability of visual cues using four levels of fog (from a completely clear view to impenetrable fog where no visual cues to the targets position were available) as well as the presence of sound cues emitted from the object (sound present/absent). The maximum correlation between stimulus and response was higher when sound was present in some conditions, particularly when visual uncertainty was high, while the tracking delay remained unaffected across all fog levels. Using a Kalman filter to estimate the underlying sensory noise, we found strong evidence that sensory noise was lower when sound was present than when sound was absent both for manual and for ocular tracking, particularly for those conditions with higher visual uncertainty. In exploratory analyses, we further show strong correlations between manual and ocular tracking in all measures (maximum correlation, tracking delay, sensory precision). However, when isolating the multisensory advantage, these correlations all but disappeared for maximum correlation and tracking delay, while remaining substantial for sensory precision. Similarly, behavioral tracking correlated generally strongly with underlying sensory noise, but much less so when it came to the advantage conferred by added sound cues. Our results show that continuous psychophysics is well-suited for the study of multisensory integration, particularly when a Kalman filter analysis is used to estimate sensory uncertainty from behavioral data.
Wexler, M.
Show abstract
Recent work has brought to light a number of stimulus families whose perception is shaped by strong idiosyncratic biases. These biases differ significantly from one observer to the next, yet remain quite stable within observers when measured over multiple points in time, sometimes over months or even years. Nevertheless, we have previously shown that at least some of these biases undergo small but systematic changes over time. Although these temporal changes are also idiosyncratic, they generally act as a kind of memory that accumulates small random steps. Other research has shown that when stimuli are shown at different points in the visual field, biases can vary idiosyncratically across the spatial field as well. Here we ask whether variations in biases across space follow any regular pattern, and whether spatial and temporal variations are independent of one another. Measuring biases for surface orientation in structure-from-motion stimuli, sampled at numerous points in space and time, we find that variations both in time and space are positively autocorrelated: the closer two points are to each other in space or in time, the more similar the biases at those two points. We also found that spatial and temporal variations of bias are correlated, both between and within participants. Bias variations over space, time, and space-time are therefore not random but follow dynamics that may provide clues about the underlying mechanisms.
Xue, S.; Landy, M.; Carrasco, M.
Show abstract
In human adults, visual performance varies systematically around the visual field. It is higher along the horizontal than the vertical meridian (horizontal-vertical anisotropy, HVA) and higher at the lower than the upper vertical meridian (vertical-meridian asymmetry, VMA). Although these robust performance fields have been linked to non-uniform neural resources, the system-level computations that translate neural constraints into perceptual asymmetries remain largely unexplored. Here, we used reverse correlation to characterize feature weighting and internal noise during peripheral orientation detection. Reverse correlation revealed non-ideal feature weighting in the joint orientation-spatial-frequency space, which was incorporated into a noisy-observer model jointly constrained by trial-wise detection responses and double-pass consistency. Across observers, the magnitude of the HVA in contrast sensitivity was correlated with individual asymmetries in orientation sensitivity and additive internal noise. In contrast, we found limited evidence that any tested representational or noise components reliably accounted for individual differences in VMA magnitude. Spatial-frequency tuning exhibited substantial individual variability, often peaking below the signals spatial frequency, but did not vary systematically across locations or explain performance asymmetries. These findings suggest that the HVA reflects systematic variation in the feature weighting of task-relevant orientations and internal noise, while constraining which computational components provide robust explanations of polar-angle asymmetries. Moreover, our framework provides a principled approach that links this prevalent perceptual asymmetry to the system-level computations that transform sensory information into perceptual decisions. Author summaryHuman vision is surprisingly uneven across our field of view. At the same distance from where we look, vision is better along the horizontal axis than the vertical axis, and better in the lower than the upper half of the vertical axis. These differences are well established, but little is known about their computational basis. To investigate, we combined visual detection tasks with computational modeling. By analyzing how observers detected faint patterns within noisy images, we measured how they process distinct features--like patterns with different orientations and spatial frequencies at different visual field locations. We then fit a mathematical model to separate distinct sources of internal noise. We found that differences in how the brain processes line orientations predicted the magnitude of the horizontal-vertical asymmetry across individuals, whereas distinct components of internal noise predicted these asymmetries in different ways. Our results show that visual field asymmetries are not driven by a single visual bottleneck, but rather by location-specific combinations of how the brain encodes relevant feature information and neural noise. This study helps explain why human vision is fundamentally uneven across our field of view.
Razafindrahaba, A.; Koiso, K.; van de Ven, V.; De Martino, F.; De Weerd, P.; Roberts, M. J.
Show abstract
Filling-in occurs during the perceptual disappearance of a blank figure presented on a textured background. Current models of perceptual filling-in are based on a two-stage model where the figure boundary weakens after a period of adaptation, followed by the spreading of the background representation into the region representing the figure. This suggests a competition between figure boundary and background representations whereby filling-in is facilitated by a weaker boundary representation and a stronger background representation. Here, we test this interpretation, by using the oblique effect and surround-modulation suppression, which are functional properties of early visual cortex that modulate the expected strengths of the responses to the background texture and to the figure boundary. In a sample of N=58 participants, we found more filling-in with background textures of cardinal compared to oblique orientations (earlier onset time, with more and longer episodes of filling-in per trial), in line with a known, stronger neuronal response for cardinal than for oblique orientation in early visual cortex. We found more filling-in when the main axis of the rectangular figure was iso-oriented rather than cross-oriented with the background texture (more and longer episodes of filling-in per trial, but no change in onset time), in line with a lower response to oriented stimuli when surrounded by iso-oriented flankers compared to cross-oriented flankers. Overall, our results support the two-stage model and suggest the involvement of early visual cortical areas characterized by the oblique effect and orientation- tuned surround-suppression.
Pandey, A.; Nadeem, A.; Harris, L. R.; Jörges, B.
Show abstract
During sideways movement of an observer, optic flow parsing - in which an objects speed in the world is extracted from all the other visual movement present in the scene, self-generated and otherwise - has been shown to be incomplete, leading to biases in speed perception, particularly when object and observer are moving in opposite directions. Here, we assess how judgements about the speed of objects moving in depth (judged relative to the world) towards or away from an observer (6 m/s) are affected by simultaneous movement of the observer either in the same or opposite direction as the object. In a virtual reality display, participants (n = 25) viewed a sphere simulated as moving in a corridor either while they were stationary or during visually simulated self-motion in the same or opposite direction as the object. They judged the spheres movement relative to the world by comparing its motion to a probe sphere that travelled laterally across the corridor in front of them. In a second experiment (n = 28) participants performed the same task but during faster self-motion (10 m/s). The second cohort also judged the direction in which the object was perceived to move during the same combinations of self and object speeds. Object speed was overestimated when the object travelled in the direction opposite to the observer compared to how objects motion was judged when the observer was stationary. However, object speed was also overestimated during self-motion in the same direction as the object where participants were also much more likely to misjudge the direction of motion of the object. Precision of judgements was lower when self-motion was simulated than it was for stationary observers. A simple arithmetic model of flow parsing fails to capture these results satisfactorily, suggesting that different mechanisms may be at play when the observer travels in the same direction as a moving object and is vulnerable to misperceiving its direction of travel.
Kling, S. M.; Lascombes, U.; Nau, M.; Masson, G. S.; Szinte, M.
Show abstract
Eye movements provide valuable insights into human cognition and are a critical variable in numerous functional magnetic resonance imaging (fMRI) studies. Yet, when the eyes are closed, camera-based eye-tracking is unavailable, making studies of eyes-closed states challenging. Here, we address this gap through MR-based gaze decoding with DeepMReye, a deep learning framework for camera-free reconstruction of gaze behavior from the MR-signal of the eyes. We first show that fine-tuning DeepMReye using visuomotor calibration data acquired when the eyes were open significantly improves gaze decoding, and that this fine-tuning does not require simultaneous camera-based data. We next assessed whether model performance could be further improved by incorporating data acquired while participants gazed at known positions with both eyes open and closed. Notably, while DeepMReye was originally trained exclusively on eyes-open data, the network successfully generalized eyes-closed periods, with performance improving significantly through fine-tuning on the eyes-closed data. These findings demonstrate that reliable gaze monitoring during eyes-closed periods is feasible, enabling a more effective integration of eye-tracking in fMRI research and, consequently, advancing our understanding of human cognition.